Attention, Perception, & Psychophysics
○ Springer Science and Business Media LLC
Preprints posted in the last 30 days, ranked by how well they match Attention, Perception, & Psychophysics's content profile, based on 17 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Chen, S.; Mueller, H. J.; Shi, Z.
Show abstract
Attentional control balances proactive suppression of predictable distractors with reactive suppression of unexpected ones. Yet, how internal states such as alertness shape this balance is unclear. Using pupillometry and eye tracking across two probability-cueing experiments (conducted in 2024) with varying distractor prevalence, we distinguished tonic (baseline pupil size across blocks) from trial-level pupil size fluctuations (trial-by-trial residual variability in pre-stimulus pupil size). With moderate prevalence, suppression of frequent-region distractors developed gradually, whereas high prevalence induced near-immediate suppression. Behavioral measures (e.g., reaction times) were closely linked to tonic and trial-level pupil size fluctuations. Critically, both alertness components jointly influenced control: during early learning, heightened trial-level pupil size increased distractor capture and reduced target fixations, whereas later on, suppression shifted to a proactive mode resilient to trial-level fluctuations. Under high prevalence, this shift occurred faster. Notably, higher trial-level pupil size generally accelerated first target selection. These findings show that tonic alertness and trial-level alertness fluctuations dynamically regulate reactive and proactive control during statistical learning. Impact StatementThis study shows that people become better at ignoring predictable distractions over time, but that this improvement depends not only on what they have learned about the task environment, but also on their current level of alertness. By combining eye tracking and pupil measures, we found that temporary increases in alertness can sometimes help people orient more quickly to relevant information, yet during earlier stages of learning they can also make attention more vulnerable to distracting events. These findings suggest that successful focus in complex environments depends on a dynamic interplay between learned expectations and moment-to-moment fluctuations in mental state, with implications for understanding sustained attention in settings such as monitoring, driving, and other tasks that require people to stay engaged while resisting distraction.
Andrade, K. D.; Melton, D. L.; Ries, S. K.
Show abstract
Language production requires the coordination of multiple cognitive processes. The ability to anticipate and override a habitual response in favor of a contextually-appropriate response are key subprocesses of cognitive control which enable speakers to communicate effectively. Word retrieval involves the co-activation of semantically related alternatives from which the speaker must select the appropriate target representation. Although cognitive control mechanisms have been proposed to contribute to resolving semantic interference during language production, the nature of these control processes remain unclear. Studies investigating the temporal dynamics of cognitive control during decision making tasks have led to a distinction between two operating processes: proactive control, initiated prior to the occurrence of conflict, and reactive control recruited after conflict is detected. We investigated the roles of proactive and reactive control in resolving interference between competing linguistic representations during word retrieval. We analyzed congruency sequence effects combined with delta-plot distributional analyses to dissociate potential adjustments in proactive versus reactive cognitive control in a picture-naming task manipulating semantic context compared to a minimally-linguistic Stroop-like paradigm. Reaction time distributional properties following semantically related trials revealed the engagement of proactive control in semantic interference resolution during word retrieval in the PWI task. In contrast, reactive inhibitory control was engaged in resolving semantic interference following low conflict trials. This distinction was not present in the minimally-linguistic task, which did not appear to engage adaptive control to the same extent. These findings demonstrate that both proactive and reactive cognitive control mechanisms contribute to language production, and are engaged dynamically, adjusting trial-by-trial to resolve semantic interference during word retrieval. In addition, our study provides important insight into the comparison of language with other cognitive domains and positions linguistic paradigms as being instrumental in the study of cognitive control dynamics.
Zimmermann Bortoluzzi, L.; Rohenkohl, G.
Show abstract
During active vision, the brain must coordinate where to move the eyes with predictions about upcoming sensory input. Before each saccade, perception is enhanced at the upcoming fixation location, but whether this enhancement depends on expectations about target features remains unknown. Here, participants prepared a saccade to a cued location while reporting the presence and orientation of a brief visual target that appeared either at the saccade goal or at the opposite location. Feature expectation was manipulated across blocks by varying the probability of the two target orientations. Perceptual sensitivity (d') increased when targets were presented at the saccade goal, consistent with presaccadic enhancement, and was also higher for less expected features. However, these effects were independent: feature probability did not alter the magnitude of presaccadic enhancement. Moreover, presaccadic enhancement increased near saccade onset, whereas the advantage for less expected features weakened as movement onset approached. Saccade latency revealed a contrasting pattern. Visual targets presented at the saccade goal delayed movement initiation. This delay depended on feature probability, with longer latencies for unexpected than for expected features only when saccades were directed towards the target. This location-specific effect persisted after accounting for perceptual report, and the latency cost for unexpected features was reproduced in a follow-up experiment. Together, these findings show that feature probability enhanced sensitivity to unexpected information independently of presaccadic enhancement, while selectively delaying saccade initiation towards targets with unexpected features. This dissociation suggests that feature expectation modulates perception and action through functionally distinct forms of visual processing.
Callahan-Flintoft, C.; Larkin, G. B.
Show abstract
Visual search is a critical component of many professions such as military operations, baggage screening, and radiology. Aided Target Recognition (AiTR) systems are designed to highlight potential threats across the operator visual field in real-time, directing attention and improving accuracy. However, these systems may impact search and, consequently, situational awareness by diverting attentional resources from non-highlighted, yet relevant, locations. Previous work suggests that scene gist is extracted within the first 250 ms of scene onset (Vo & Henderson, 2010). As such, this study examined whether a 250 ms AiTR onset delay could encourage a more even distribution of attention. Participants searched synthetically generated scenes and classified each person in the scene as armed or unarmed. Depending on their condition, participants either saw the scenes unaugmented (No AiTR condition), with AiTR highlights consisting of red bounding boxes around armed people and yellow boxes around unarmed (AiTR condition), or with AiTR highlights presented 250 ms post scene onset (Delayed AiTR condition). A surprise memory test of background objects presented in the search scenes was administered to all participants upon completion of the search task. As predicted and preregistered, results showed less overt attentional deployment to background information (anything other than the people themselves) in the AiTR condition compared to No AiTR , however, decreased overt attentional deployment was not seen in the Delayed AiTR group. A similar pattern was observed in the memory data (with the AiTR condition having a lower score than the No AiTR condition and the Delayed AiTR condition), this difference was not significant.
Herrmann, B.; Fink, L. K.; Pandey, P. R.; Johnsrude, I.; Ryan, J. D.
Show abstract
Speech comprehension in noisy environments often requires cognitive effort, but listeners may disengage when comprehension becomes impossible. Eye movements have recently emerged as a promising new measure of listening effort, but it remains unclear whether eye movements are sensitive to the full effort profile across easy, difficult, and impossible speech comprehension. Across four experiments, participants listened to sentences at easy, difficult, and impossible levels of multi-talker background babble while pupil size and eye movements were recorded. Pupil size generally followed the expected inverted u-shaped effort profile: low for easy speech, maximal for difficult but still intelligible speech and lower again for impossible speech, although this pattern partly reflected sustained, condition-specific differences and not only sentence-evoked responses. Gaze dispersion - measuring the spread of eye movements - decreased with high temporal selectivity during difficult relative to easy and impossible speech, indicating reduced eye movements during active, effortful listening. However, gaze dispersion was also lower, but less temporally selective, during impossible compared to easy listening, especially in non-baseline-corrected analyses, suggesting that reduced eye movements do not index listening effort uniquely. Instead, eye movements appear to reflect both attentional engagement during difficult listening and disengagement or inward attention when meaningful listening is no longer possible. These findings indicate that pupil size and eye movements provide complementary indices of listening-related cognition, and highlight the integration of listening, cognition, and motor systems.
Malik, A.; Kolmel, L.; Billino, J.; Doerschner, K.
Show abstract
Humans rely on multiple sensory modalities, such as vision, audition, and touch, to perceive materials in everyday life. Previous research shows that multisensory perception leads to facilitation, yet the mechanisms responsible for this facilitation remain poorly understood. One potential mechanism is crossmodal prediction, whereby input from one modality generates predictions about another. While substantial research on multisensory facilitation has focused on bottom-up processes, such as spatial, temporal, and semantic congruency, the role of crossmodal predictions, particularly in material perception, has received little attention. To address this gap, we conducted two experiments, a reaction time task and a material rating task, in which participants viewed computer-generated animations of familiar objects being dropped to the ground. The paradigm exploited the natural temporal structure of impact events: pre-impact visual appearance provides information about an objects material and therefore can generate expectations about the forthcoming impact sound. Critically, participants saw the event only until before the impact, after which the video was masked. Thus, vision and audition were temporally aligned but not presented concurrently, allowing us to isolate the influence of visually driven expectations on the incoming auditory information without a bottom-up conflict. In some trials, the sound matched the expected material, but in a subset, it was incongruent, violating expectations elicited by the preceding visual information. Across both experiments, participants took longer to respond on incongruent than congruent trials, suggesting increased processing demands. In the rating task, incongruent trials also shifted material judgments, such that ratings reflected a weighted combination of incoming auditory information and visually driven predictions, with large individual differences in relative cue weighting. These findings suggest that priors on material properties from one modality, specifically vision, not only establish high-level expectations within the modality about an objects future state, but also extend across modalities.
Pandey, A.; Nadeem, A.; Harris, L. R.; Jörges, B.
Show abstract
During sideways movement of an observer, optic flow parsing - in which an objects speed in the world is extracted from all the other visual movement present in the scene, self-generated and otherwise - has been shown to be incomplete, leading to biases in speed perception, particularly when object and observer are moving in opposite directions. Here, we assess how judgements about the speed of objects moving in depth (judged relative to the world) towards or away from an observer (6 m/s) are affected by simultaneous movement of the observer either in the same or opposite direction as the object. In a virtual reality display, participants (n = 25) viewed a sphere simulated as moving in a corridor either while they were stationary or during visually simulated self-motion in the same or opposite direction as the object. They judged the spheres movement relative to the world by comparing its motion to a probe sphere that travelled laterally across the corridor in front of them. In a second experiment (n = 28) participants performed the same task but during faster self-motion (10 m/s). The second cohort also judged the direction in which the object was perceived to move during the same combinations of self and object speeds. Object speed was overestimated when the object travelled in the direction opposite to the observer compared to how objects motion was judged when the observer was stationary. However, object speed was also overestimated during self-motion in the same direction as the object where participants were also much more likely to misjudge the direction of motion of the object. Precision of judgements was lower when self-motion was simulated than it was for stationary observers. A simple arithmetic model of flow parsing fails to capture these results satisfactorily, suggesting that different mechanisms may be at play when the observer travels in the same direction as a moving object and is vulnerable to misperceiving its direction of travel.
Faul, F.; Nuthmann, A.
Show abstract
Current debates regarding the relative contribution of saliency versus semantics to gaze control often rely on comparing the predictive power of saliency and meaning maps. We argue that such indirect, global approaches are fundamentally limited because fixations arise from heterogeneous, local causes that are conflated in whole-scene comparisons. To substantiate this claim, we used a direct method where participants explicitly identified the reasons for fixation at specific clusters of high fixation density, distinguishing between low-level saliency and various semantic categories, as well as the most important one. The obtained judgments revealed that multiple factors contribute simultaneously to gaze control. Although their influence varied across fixation clusters, semantics generally dominated saliency. Notably, abstract semantic categories, particularly "unknown/unusual," proved important, highlighting the role of prior knowledge and novelty besides personal relevance in guiding attention. To interpret these findings in the context of existing models, we propose a framework distinguishing between processes highlighting interesting locations in the image from a sampling strategy translating this information into scanpaths. Within this framework, classic saliency and meaning maps are viewed as restricted inputs to the strategy, whereas deep learning-based models (e.g., DeepGaze IIE) are more general and may also implicitly encode aspects of the strategy itself. Consistent with this, we found that the predictive performance of DeepGaze IIE varied less significantly with the specific reasons for fixation than that of classic saliency and meaning map approaches.
Hooper, J.; Dengler, J.; Basilico, D.; Nelson, M. J.
Show abstract
Sentence comprehension requires the incremental construction of syntactic structure and semantic interpretation. Prior neural work (Nelson et al., 2017) identified key neural events at major phrase boundaries during sentence comprehension. To investigate a behavioral correlation of these processes, we used self-paced reading to examine the impact of syntactic phase boundaries, semantic congruence, and sentence structure on sentence processing. Participants read object-relative, subject-relative, and canonical control sentences one word at a time and a subsequent comprehension task. Reading times were analyzed relative to phrase boundaries, node-closing operations, and semantic congruence. Object-relative sentences produced the greatest processing difficulty, demonstrated by increased reading times and decreased comprehension accuracy. Reading times peaked at the phrase boundaries, indicating that processing costs are tied to constituent completion rather than individual lexical categories. Reading times also increased with the number of syntactic constituents completed at a phrase boundary. Agent-patient semantic congruence produced its largest effects in object-relative sentences, suggesting that semantic information interacts with syntactic computations when processing demands are greatest. These findings demonstrate that self-paced reading is sensitive to the incremental processing associated with syntactic constituent completion. Processing costs are tied more closely to phrase completion than to individual lexical categories, scale with the amount of syntactic structure completed at a boundary and interact with agent-patient semantic interpretation during object-relative sentence comprehension. Together, these findings support a view of sentence comprehension in which syntactic structure building and semantic interpretation proceed incrementally and interact continuously throughout online language processing.
Paro, A. N.; Sheikh, B. I.; Stanford, T. R.; Salinas, E.
Show abstract
The ability to orient or attend to sensory events is generally greater in response to visual and auditory cues occurring together than in response to single-modality cues occurring alone. In such cases the perceptual fusion of cross-modal stimuli (multisensory integration) depends on low-level features (e.g., location, intensity) and follows well established principles. However, less is known about multisensory integration mechanisms when behavioral responses are less direct and require top-down control. Here we investigate this in human participants using an urgent multisensory choice task that effectively dissociates stimulus-driven and goal-driven contributions to performance based on their distinct temporal signatures. Task conditions varied the modality of the cues (auditory, visual, or both), their location (left or right), and the rule defining the correct choice (look toward or away from a given cue). When spatially coincident cues were associated with the same response rule ("look away"), we observed multisensory enhancement and performance remained close to a statistical expectation as the choice process unfolded. However, when spatially disparate cues were associated with different rules but the same target, one cue dominated performance and the other produced crossmodal capture, i.e., low-level competition. The results indicate that the efficacy of multisensory integration is dictated by the stimulus-and goal-driven signals produced by each cue, with all four factors rapidly interacting in accordance to the dynamics of spatial attention. Significance StatementAuditory and visual stimuli located near each other in space and time are typically bound into a single sensory percept that draws attention most effectively. However, it is unclear whether such "multisensory integration" occurs during behaviors that go beyond directly attending or orienting to cue stimuli and require top-down control. We investigated this using a novel task design with which stimulus-driven and goal-driven contributions to performance can be accurately identified. We found that multisensory enhancement depends not so much on the complexity of the requested cue-response associations, but rather on the timing and alignment of the stimulus-and goal-driven signals derived from each cue (auditory and visual) -- similar to the way that such signals dictate the allocation of spatial attention.
Peterzell, D. H.; Arrighi, R.; Di Cesare, C.; Gurioli, M.; Farini, a.; Grasso, P. A.
Show abstract
Numerosity adaptation (the underestimation of number after exposure to a numerous adaptor) is reduced when adaptor and test differ in color, suggesting that the numerosity system parses items into color-defined categories. Here we ask whether this chromatic selectivity is organized into multiple narrowly tuned chromatic channels, and whether its expression depends on individual chromatic sensitivity. Twenty observers (aged 22-61) completed two psychophysical tasks. First, chromatic discrimination was measured for five hues spaced in 5{degrees} CIE L*a*b* steps ({Delta}H = 0{degrees}, 5{degrees}, 10{degrees}, 15{degrees}, 20{degrees}) from a red reference (LCh: 54, 118, 38), yielding an individual just-noticeable difference (JND). Second, numerosity adaptation was measured across the same five chromatic distances between a 48-dot adaptor and the test. Observers with superior discrimination (JND < 2.5{degrees}) showed robust chromatic tuning, adaptation declining as the test moved away from the adaptor hue, whereas poorer discriminators showed none. Using an interindividual-covariance / factor-analytic approach, we found that adaptation strengths at neighboring chromatic distances were highly correlated and fell off with chromatic separation. Principal component analysis extracted two factors, one loading on the larger chromatic distances and one on the smaller; under oblique (promax) rotation the two factors were substantially correlated (r = .66), implying at least two dissociable but overlapping chromatically tuned mechanisms. These results suggest that numerosity adaptation is mediated by multiple, comparatively narrow chromatic channels, resembling the higher-order color mechanisms inferred from color scaling, SSVEP, and fMRI, rather than the two early cardinal axes (L-M, S-(L+M)).
Perez, O. D.; Cancino, N.; Hermosilla, D.; Soto, F. A.; Vogel, E. H.
Show abstract
In animal learning research, learning is often represented by plotting a behavioral measure as a function of training trials. A particularly clear case is habituation, a basic form of learning in which repeated presentation of a stimulus produces a decrement in responding. Although retention tests provide the strongest basis for evaluating durable habituation once short-lived performance effects have dissipated, the pattern of response change across stimulus repetitions, or habituation curve, remains theoretically and empirically relevant because it is used to characterize determinants of habituation, individual and clinical profiles, and functional forms, including linear, curvilinear, asymptotic, and mixed incremental-decremental patterns of responding. However, group averaged curves may conceal substantial individual heterogeneity. Here, we analyzed archived human eyeblink habituation data from 157 participants to ask whether the curve shape selected for the group average reflects the curve shapes observed at the individual level. Five candidate functions were fitted separately to each participant and to the corresponding group average. No single function characterized most individuals. More importantly, the model selected for the group average differed from the most frequent individual model in all four groups. When data were pooled across groups, the average favored a dual-process form, a shape that matched the individual plurality in none of them. Simulation analyses showed that averaging heterogeneous individual trajectories can itself produce a group curve that favors a more complex model. Our findings show that group averaged habituation curves should not be treated as direct descriptions of the typical individual trajectory.
Dirks, C. E.; Guest, D. R.; Oxenham, A.
Show abstract
Context effects are ubiquitous across sensory systems and reflect a general encoding principle for both simple and complex stimuli. One simple context effect, contraction bias, manifests in two-interval perception tasks as a bias of the perceived magnitude of the first stimulus toward the center of the overall magnitude range. The underlying cause of contraction bias is unclear. One explanation is that a listeners magnitude estimate of the first stimulus is combined with a perceptual anchor, usually the mean stimulus magnitude, biasing it toward the anchor (sensory model). An alternative explanation is that a listeners response criterion shifts, based on the magnitude of the stimulus pair, relative to the mean magnitude of the stimuli range (decision model). Two pitch-discrimination experiments were performed to test these hypotheses in the auditory domain. The first was a forced-choice discrimination task, where listeners were asked to identify the higher or lower tone in a pair. The second was a same-different task where listeners indicated whether or not the two tones in a pair differed in frequency. Contraction bias was observed in the higher-lower discrimination task, even after extensive perceptual training with feedback. In contrast, no contraction bias was observed in the same-different task. Computational models of the sensory and decision hypotheses were fit to data from both experiments. The sensory model captured the pattern of results the higher-lower experiment but erroneously predicted a contraction bias in the same-different task. The decision model produced similar predictions to the sensory model in the higher-lower task but correctly predicted no contraction bias in the same-different task, and produced lower prediction errors and more stable parameter estimates in both paradigms. Overall, the results suggest that the underlying nature of the contraction bias may reflect decision, rather than sensory, biases based on the context.
Husta, C.; Seijdel, N.; Drijvers, L.
Show abstract
Face-to-face communication requires listeners to attend, integrate, and weigh multiple communicative signals, including auditory speech, mouth movements, and co-speech gestures. The contribution of these signals may depend on the reliability of auditory input and the informativeness of the available signals. We utilized rapid invisible frequency tagging (RIFT) with EEG to examine how participants attend to and integrate these different signals in clear and adverse listening conditions. Participants watched videos of an actress producing clear or noise-vocoded sentences. Auditory speech was amplitude-modulated at 58Hz, while the luminance of the gesture and mouth regions was frequency-tagged at 63Hz and 65Hz. Degraded speech elicited stronger responses at the auditory tagged frequency, suggesting increased attentional gain to the auditory signal when listening was challenging. In contrast, clear speech elicited stronger responses at the gesture tagged frequency and a stronger 2Hz intermodulation response (65-63Hz), reflecting enhanced nonlinear coupling between mouth movements and gestures. Finally, in degraded speech, the informativeness of mouth movement, but not gesture, was associated with intermodulation strength, suggesting that the informativeness of mouth movements plays a greater role in multisensory interaction when listening is challenging. Our findings demonstrate that both signal reliability and informativeness shape multisensory integration during spoken language comprehension.
Pesthy, O.; Toth-Faber, E.; Nagy, C.; Nemeth, M.; Janacsek, K.; Nemeth, D.
Show abstract
Children often outperform adults in probabilistic statistical learning tasks, yet the mechanisms underlying this developmental advantage remain poorly understood. Here, we used eye-tracking measures of belief updating to examine how children and adults acquire and update predictions in a probabilistic sequence-learning task. Using the standard (oculomotor) reaction time measure, children showed stronger statistical learning than adults, replicating previous behavioral findings while revealing a more detailed profile of developmental differences in statistical learning. Critically, children updated their predictions more frequently: they were less likely to repeat previous predictions and more likely to shift their expectations in response to new input. Adults, in contrast, showed greater persistence, tending to maintain prior predictions even when those predictions were inconsistent with the underlying statistical structure. Despite these pronounced differences in updating behavior, the processing and use of prediction errors were remarkably similar across age groups. These findings indicate that developmental differences in statistical learning do not primarily arise from how prediction errors are computed, but rather from how prior beliefs and incoming information are weighted during belief updating. Children's enhanced learning may therefore reflect reduced reliance on stable priors and greater sensitivity to current sensory evidence, supporting a more exploratory learning strategy. Adults, by contrast, appear to favor an exploitative strategy that stabilizes existing predictions but reduces flexibility in probabilistic environments. More broadly, the results suggest that developmental changes in statistical learning may reflect age-related differences in how readily learners revise their predictions in response to incoming evidence. By integrating sensitive oculomotor measures with analyses that probe the mechanisms underlying belief updating, the present study provides a more fine-grained account of how predictive learning changes across development and offers a framework for reconciling previously inconsistent developmental findings in statistical learning.
XU, M.; REN, Y.
Show abstract
Building upon foundational psychological theories of event segmentation, this study addresses the limitation of overreliance on temporal boundaries as the primary segmentation criterion. Drawing on two experiments of direct and indirect causation in Mandarin Chinese, this study demonstrates how cognitive segmentation granularity and semantic integration jointly shape syntactic encoding. Results reveal distinct event encoding patterns for direct and indirect causation: coarse-grained segmentation leads to compact syntactic structures (e.g., verb-resultatives), while fine-grained segmentation yields varied multi-clausal expressions. Chinese speakers update event models via prediction errors of intentionality and protagonists, and tend to establish event boundaries at goal-relevant action endpoints when construing causal chains. These conceptual dimensions exert a modulating influence on both event segmentation and semantic integration. We propose a triad model integrating event segmentation, semantic integration, and linguistic specificity, providing a unified framework for elucidating the mind-language interface in conceptual construction and event coding of causation.
Makhsous, M.; Jowkar, M.; Rezayat, E.
Show abstract
Studying chess experts helps researchers understand how intensive practice shapes thinking skills. Cognitive flexibility is the ability to adjust thoughts when rules or tasks change. Working memory is the ability to hold and use information over short periods. This study compared cognitive flexibility and working memory precision between adolescent chess players and non-players. Twenty-four professional chess players and twenty-five controls completed two novel behavioral tasks. Chess players showed better accuracy in both tasks than controls. They adapted more efficiently when rules changed during a continuous learning task. They also remembered facial expressions more precisely in a working memory task. Learning rates in the flexibility task did not differ between groups. These results indicate that chess expertise may improve rule-guided flexibility and visual working memory precision in adolescents.
Koenigsmark, V. T.; Stenner, M.-P.; Reeder, R. R.; Azanon, E.
Show abstract
The absence of voluntary visual imagery, known as aphantasia, offers a unique lens into the role of visual imagery in visual memory processes such as face recognition. While aphantasics often report difficulties, behavioral differences in standard tasks have generally been small. One possibility is that the contribution of visual imagery becomes apparent only when face recognition is especially demanding. We compared age- and gender-matched aphantasics and typical imagers on the more challenging long-form version of the Cambridge Face Memory Test (CFMT+) and on a measure of inverted face recognition. We also included tasks assessing object recognition and face perception. No group differences emerged for face perception or object recognition. By contrast, typical imagers outperformed aphantasics under high visual and mnemonic demands in face recognition in the laboratory cohort, particularly at the highest difficulty level of the CFMT+ and in inverted face recognition. This effect was attenuated or even absent in the online cohort. Drift-diffusion modelling indicated that this discrepancy was primarily driven by reduced response caution in online typical imagers. A meta-analysis of published short-form CFMT studies (total N = 432) revealed a moderate and reliable advantage for typical imagers (hedges g [≤] 0.41). Finally, across cohorts and tasks, aphantasics reported consistently lower subjective confidence, independent of accuracy. Overall, these findings suggest that visual imagery benefits face recognition, and highlight the need for caution in online testing, the predominant approach in aphantasia research.
Razafindrahaba, A.; Koiso, K.; van de Ven, V.; De Martino, F.; De Weerd, P.; Roberts, M. J.
Show abstract
Filling-in occurs during the perceptual disappearance of a blank figure presented on a textured background. Current models of perceptual filling-in are based on a two-stage model where the figure boundary weakens after a period of adaptation, followed by the spreading of the background representation into the region representing the figure. This suggests a competition between figure boundary and background representations whereby filling-in is facilitated by a weaker boundary representation and a stronger background representation. Here, we test this interpretation, by using the oblique effect and surround-modulation suppression, which are functional properties of early visual cortex that modulate the expected strengths of the responses to the background texture and to the figure boundary. In a sample of N=58 participants, we found more filling-in with background textures of cardinal compared to oblique orientations (earlier onset time, with more and longer episodes of filling-in per trial), in line with a known, stronger neuronal response for cardinal than for oblique orientation in early visual cortex. We found more filling-in when the main axis of the rectangular figure was iso-oriented rather than cross-oriented with the background texture (more and longer episodes of filling-in per trial, but no change in onset time), in line with a lower response to oriented stimuli when surrounded by iso-oriented flankers compared to cross-oriented flankers. Overall, our results support the two-stage model and suggest the involvement of early visual cortical areas characterized by the oblique effect and orientation- tuned surround-suppression.
Stewart, E. E. M.; Wagner, I.; Schuetz, A. C.; Fleming, R. W.
Show abstract
The ability to mentally rotate objects is a fundamental feature of human cognition, and humans can use this ability to make choices about objects based on their geometry. However, remarkably little is known about how such choices are reached, and what sort of visual information might facilitate them. We devised an experiment where participants had to mentally simulate an object's rotation to choose which of two objects was better for a subsequent task based on its shape alone. We also tracked their gaze while they made their choice, to see which visual information they were using to facilitate this mental simulation. We found that participants were consistently able to choose the most suitable object for the task, and, remarkably, the visual information they sampled was directly linked to their choices. Put simply, participants made better choices when they looked at more informative regions of the objects, and participants who sampled regions that were better for facilitating mental simulation made better choices overall. These findings reveal a direct link between fixations, simulation, and decision-making, suggesting that to perform any fine-grained mental simulation people need to direct their gaze at specific, informative points of an object to simulate its two-dimensional proximal image displacement.